Back

Genomics, Proteomics & Bioinformatics

All preprints, ranked by how well they match Genomics, Proteomics & Bioinformatics's content profile, based on 188 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
circASbase:A Comprehensive Database of Alternative Splicing Events in circRNAs

Zou, L.; Jian, Z.; Li, H.; Xu, C.; Wang, Y.; Guo, X.; Song, X.

2025-03-13 bioinformatics 10.1101/2025.03.09.642279 medRxiv
Top 0.1%
81.0%
Show abstract

Despite extensive studies highlight the critical roles of alternative splicing in generating mature circRNA isoforms and enhancing their function diversity, a significant gap remains in the availability of dedicated databases for circRNA alternative splicing events. To bridge this gap, we developed circASbase, a pioneering and comprehensive database that catalogues 452,129 alternative splicing events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNAs. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of alternative splicing events, highlighting the unique regulatory landscape of circRNAs. These unique splicing events result in functional differences of circRNAs by affecting IRES sites, m6A sites, ORFs, protein features, miRNA targets, and more. In summary, circASbase not only covers the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of alternative splicing events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/

2
Characterization of RNA editing profiles in rice endosperm development

Chen, M.; Xia, L.; Tan, X.; Gao, S.; Wang, S.; Li, M.; Zhang, Y.; Xu, T.; Cheng, Y.; Chu, Y.; Hu, S.; Wu, S.; Zhang, Z.

2024-01-30 bioinformatics 10.1101/2024.01.27.577525 medRxiv
Top 0.1%
73.4%
Show abstract

Rice (Oryza sativa L.) endosperm provides nutrients for seed germination and determines grain yield. RNA editing, a post-transcriptional modification essential for plant development, unfortunately, is not fully characterized during rice endosperm development. Here, we conduct genome re-sequencing and RNA sequencing for rice endosperms across five successive developmental stages and perform systematic analyses to characterize RNA editing profiles during rice endosperm development. We find that the majority of their editing sites are C-to-U CDS-recoding in mitochondria, leading to increased hydrophobic amino acids, and affecting structures and functions of mitochondrial proteins. Comparative analysis of RNA editing profiles across the five developmental stages reveals that CDS-recoding sites present higher editing frequencies with lower variabilities, and recoded amino acids, particularly caused by these sites with higher editing frequencies, tend to exhibit stronger evolutionary conservation across many land plants. Based on these results, we further classify mitochondrial genes into three groups that present distinct patterns in terms of editing frequency and variability of CDS-recoding sites. Besides, we identify a series of P- and PLS-class pentatricopeptide repeat (PPR) proteins with editing potential and construct PPR-RNA binding profiles, yielding candidate PPR editing factors related to rice endosperm development. Taken together, our findings provide valuable insights for deciphering fundamental mechanisms of rice endosperm development underlying RNA editing machinery. Author summaryRice endosperm development, a critical process determining quality and yield of our mankinds essential food, is regulated by RNA editing that provokes RNA base alterations by protein factors. However, our understanding of this regulation is incomplete. Hence, we systematically characterize RNA editing profiles during rice endosperm development. We find that editing sites resulting in amino acid changes, called "CDS-recoding", predominate in mitochondria, leading to increased hydrophobic amino acids and affecting structures and functions of proteins. Comparative analysis of RNA editing profiles during rice endosperm development reveals that CDS-recoding sites present higher editing frequencies with lower variabilities. Furthermore, evolutionary conservation of recoded amino acids caused by these CDS-recoding sites is positively correlated with editing frequency across many land plants. We classify mitochondrial genes into three groups that present distinct patterns in terms of editing frequency and variability of CDS-recoding sites, indicating different effects of these genes on rice endosperm development. In addition, we identify candidate protein factors associated closely with RNA editing regulation. To sum up, our findings provide valuable insights for fully understanding the role of RNA editing during rice endosperm development.

3
A multi-tissue developmental gene expression atlas towards understanding the biological basis of phenotypes in sheep

Zhao, B.; Luo, H.; Fu, X.; Zhang, G.; Clark, E. L.; Wang, F.; Dalrymple, B. P.; Oddy, V. H.; Vercoe, P. E.; Wu, C.; Liu, G. E.; Li, C.-j.; Xiang, R.; Tian, K.; Zhang, Y.; Fang, L.

2024-11-14 genetics 10.1101/2024.11.14.623505 medRxiv
Top 0.1%
73.0%
Show abstract

Sheep (Ovis aries) represents one of the most important livestock species for animal protein and wool production worldwide. However, little is known about the genetic and biological basis of ovine phenotypes, particularly for those of high economic value and environmental impact. Here, by generating and integrating 1,413 RNA-seq samples from 51 distinct tissues across 14 developmental time points, representing early prenatal, late prenatal, neonate, lamb, juvenile, adult, and elderly stages, we built a high-resolution developmental Gene Expression Atlas (dGEA) in sheep. We observed dynamic patterns of gene expression and regulatory networks across tissues and developmental stages. When harnessing this resource for interpreting genomic associations of 48 monogenetic and 12 complex traits in sheep, we found that genes upregulated at prenatal developmental stages played more important roles in shaping these phenotypes than those upregulated at postnatal stages. For instance, genetic associations of crimp number, mean staple length (MSL), and individual birth weight were significantly enriched in the prenatal rather than postnatal skin and immune tissues. By comprehensively integrating fine-mapping results and the sheep dGEA, we identified several key genes associated with complex traits in sheep, such as SOX9 (associated with MSL), GNRHR (associated with litter size at birth), and PRKDC (associated with live weight). These results provide novel insights into the gene regulatory and developmental architecture underlying ovine phenotypes. The dGEA (https://sheepdgea.njau.edu.cn/) will serve as an invaluable resource for sheep developmental biology, genetics, genomics, and selective breeding.

4
Genetic basis of expression and splicing underlying spike architecture in wheat (Triticum aestivum L.)

Yang, G.; Pan, Y.; Cui, L.; Chen, M.; Zeng, Q.; Pan, W.; Zhe, L.; Edwards, D.; Batley, J.; Han, D.; Deng, P.; Yu, H.; Henry, R. J.; Song, W.; Nie, X.

2023-05-05 plant biology 10.1101/2023.05.04.539218 medRxiv
Top 0.1%
66.2%
Show abstract

IntroductionWheat is one of the most important staple crops worldwide, and an important source of human protein and mineral element intake. Continuously increasing stable production of wheat is critical for global food security under the challenge of population growth and limited resource input. ObjectiveSpike architecture determines the potential grain yield of wheat. However, the mechanisms of transcriptional regulation of spike architecture in wheat remain largely unknown, limiting further genetic improvement of wheat yield. In this study we explored the genetic basis of spike architecture in wheat. MethodsPopulation RNA-seq methods were used to identify the eQTLs and sQTLs associated with spike architecture and applied this to dissection of the genetic basis of gene expression and splicing controlling these complex yield-related traits. ResultsIn total, 4,143 expression quantitative trait loci (eQTLs) and 12,933 splice QTLs (sQTLs) were identified in wheat based on 178 RNA-seq samples, revealing 774 cis-eQTLs and 321 cis-sQTLs for 86 eGenes and 73 sGenes, respectively. Integration of eQTLs and sQTLs with genome-wide association study (GWAS) identified dozens of additional novel candidate genes that may contribute to spike-related traits. Gene network analysis showed that eQTLs and sQTLs were widely involved in the co-expression modules that regulate wheat spike architecture. Notably, the eQTL locus AX-108754757 regulated the expression of 5 eGenes that negatively controled grain number per spike. AX-111592099 regulated both the splicing and expression of TraesCS7B02G442100, encoding an E3 ubiquitin ligase, and playing a central role in regulating spike length. ConclusionThis study provides new insights into the genetic basis of spike architecture. This improved understanding of spike-related traits in wheat will contribute to more rapid genetic improvement.

5
A practical framework RNMF for the potential mechanism of cancer progression with the analysis of genes cumulative contribution abundance

Li, Z.; Liang, H.; Zhang, S.; Luo, W.

2021-08-13 bioinformatics 10.1101/2021.08.12.456096 medRxiv
Top 0.1%
65.9%
Show abstract

Mutational signatures can reveal the mechanism of tumorigenesis. We developed the RNMF software for mutational signatures analysis, including a key model of cumulative contribution abundance (CCA) which was designed to highlight the association between genes and mutational signatures. Applied it to 1073 esophageal squamous cell carcinoma (ESCC) and found that APOBEC signatures (SBS2* and SBS13*) mediated the occurrence of PIK3CA E545k mutation. Furthermore, we found that age signature is strongly linked to the TP53 R342* mutation. In addition, the CCA matrix image data of genes in the signatures New, SBS3* and SBS17b* were helpful for the preliminary evaluation of shortened survival outcome. In a word, RNMF can successfully achieve the correlation analysis of genes and mutational signatures, proving a strong theoretical basis for the study of tumor occurrence and development mechanism and clinical adjuvant medicine.

6
CellClick: an interactive platform for adjustable and accurate cell type annotation in single-cell and spatial omics data

Shi, L.; Dai, M.; Zhang, Y.-b.; Wu, S.; Wang, M.; Wang, X.-j.

2026-06-03 bioinformatics 10.64898/2026.06.01.727775 medRxiv
Top 0.1%
65.8%
Show abstract

Single-cell omics and spatial omics technologies are nowadays widely used in biological and medical research. In both single-cell and spatial omics data analysis, accurate cell type annotation is a key step for downstream analysis and scientific discoveries. However, high-quality cell annotation usually requires multiple rounds of manual analysis for result refinement, which poses great challenges to most researchers. Here, we present CellClick, an interactive platform for convenient and accurate cell type annotation in single-cell and spatial omics data. CellClick provides Data Preprocessing, Data Visualization, Cell Annotation, Annotation Validation, and Cell Reannotation modules, which facilitate automatic or user-guided cell selection and annotation. The feasibility of using CellClick to generate more accurate cell annotation results was exemplified by both scRNA-seq and spatial transcriptomics data.

7
Genome charaterization based on the Spike-614 and NS8-84 loci of SARS-CoV-2 reveals two major onsets of the COVID-19 pandemic

Zhang, J.; Hu, X.; Mu, Y.; Deng, R.; Yi, G.; Yao, L.

2022-12-06 bioinformatics 10.1101/2022.12.04.519037 medRxiv
Top 0.1%
62.6%
Show abstract

The global COVID-19 pandemic has lasted for three years since its outbreak, however its origin is still unknown. Here, we analyzed the genotypes of 3.14 million SARS-CoV-2 genomes based on the amino acid 614 of the Spke (S) and the amino acid 48 of NS8 (nonstructural protein 8), and identified 16 linkage haplotypes. The GL haplotype (S_614G and NS8_48L) was the major haplotype driving the global pandemic and accounted for 99.2% of the sequenced genomes, while the DL haplotype (S_614D and NS8_48L) caused the pandemic in China in the spring of 2020 and accounted for approximately 60% of the genomes in China and 0.45% of the global genomes. The GS (S_614G and NS8_48S), DS (S_614D and NS8_48S) and NS (S_614N and NS8_48S) haplotypes accounted for 0.26%, 0.06%, and 0.0067% of the genomes, respectively. The main evolutionary trajectory of SARS-CoV-2 is DS[->]DL[->]GL, whereas the other haplotypes are minor byproducts in the evolution. Surprisingly, the newest haplotype GL had the oldest time of most recent common ancestor (tMRCA), which was May 1 2019 by mean, while the oldest haplotype had the newest tMRCA with a mean of October 17, indicating that the ancestral strains that gave birth to GL had been extinct and replaced by the more adapted newcomer at the place of its origin, just like the sequential rise and fall of the delta and omicron variants. However, they arrived and evolved into toxic strains and ignited a pandemic in China where the GL strains did not exist at the end of 2019. The GL strains had spread all over the world before they were discovered, and ignited the global pandemic, which had not been noticed until the pandemic was declared in China. However, the GL haplotype had little influence in China during the early phase of the pandemic due to its late arrival as well as the strict transmission controls in China. Therefore, we propose two major onsets of the COVID-19 pandemic, one was mainly driven by the haplotype DL in China, the other was driven by the haplotype GL globally.

8
Modeling of microRNA-derived disease network repurposes methotrexate for the prevention and therapy of abdominal aortic aneurysm in mice

Shen, Y.; Gao, Y.; Shi, J.; Huang, Z.; Dai, R.; Fu, Y.; Zhou, Y.; Kong, W.; Cui, Q.

2021-12-13 bioinformatics 10.1101/2021.12.13.472366 medRxiv
Top 0.1%
62.1%
Show abstract

Abdominal aortic aneurysm (AAA) is a highly lethal vascular disease characterized by permanent dilatation of the abdominal aorta. The main purpose of the current study is to search for noninvasive medical therapies for abdominal aortic aneurysm (AAA), for which there is currently no effective drug therapy. Network medicine represents a cutting-edge technology, as analysis and modeling of disease networks can provide critical clues regarding the etiology of specific diseases and which therapeutics may be effective. Here, we proposed a novel algorithm to quantify disease relations based on a large accumulated microRNA-disease association dataset and then built a disease network that covered 15 disease classes and included 304 diseases. Analysis revealed a number of patterns for these diseases; for example, diseases tended to be clustered and coherent in the network. Surprisingly, we found that AAA showed the strongest similarity with rheumatoid arthritis and systemic lupus erythematosus, both of which are autoimmune diseases, suggesting that AAA could be one type of autoimmune disease in etiology. Based on this observation, we further hypothesized that drugs for autoimmune disease could be repurposed for the prevention and therapy of AAA. Finally, animal experiments confirmed that methotrexate, a drug for autoimmune disease, was able to prevent the formation and inhibit the development of AAA.

9
Massive Horizontal Gene Transfer in Amphioxus Illuminates the Early Evolution of Deuterostomes

Xiong, Q.; Yang, K. Y.; Zeng, X.; Wang, M.; Ng, P. K.-S.; Zhou, J.-W.; Ng, J. K.-W.; Law, C. T.-Y.; Du, Q.; Xu, K.; Falkenberg, L. J.; Mao, B.; Chen, J.-Y.; Tsui, S. K.-W.

2022-05-19 evolutionary biology 10.1101/2022.05.18.492404 medRxiv
Top 0.1%
61.5%
Show abstract

Amphioxus is considered the best-known living proxy to the chordate ancestor and an irreplaceable model organism for evolutionary studies of chordates and deuterostomes. In this study, a high-quality genome of the Beihai amphioxus, Branchiostoma belcheri beihai, was de novo assembled and annotated. Within four amphioxus genomes, twenty-eight groups of gene novelties were identified, revealing new genes that lack homologs in non-deuterostome metazoa, but share unexpectedly high similarities with those from non-metazoan species. These gene innovation events have played roles in amphioxus adaptations, including innate immunity responses, glycolysis, and regulation of calcium balance. The gene novelties related to innate immunity, such as a group of lipoxygenases and a DEAD-box helicase, boosted amphioxus immune responses. The novel genes for alcohol dehydrogenase and ferredoxin could aid in the glycolysis of amphioxus. A proximally arrayed cluster of EF-hand calcium-binding protein genes were identified to resemble those of bacteria. The copy number of this gene cluster was negatively correlated to the sea salinity of the collection region, suggesting that it may enhance their survival at different calcium concentrations. This comprehensive study collectively reveals insights into adaptive evolution of cephalochordates and provides valuable resources for research on early evolution of deuterostomes.

10
Large-scale miRNA-Target Data Analysis to Discover miRNA Co-regulation Network of Abiotic Stress Tolerance in Soybeans

Chang, H.; Zhang, T.; Zhang, H.; Su, L.; Qin, Q.-M.; Li, G.; Li, X.; Wang, L.; Zhao, T.; Zhao, E.; Zhao, H.; Liu, Y.; Stacey, G.; Xu, D.

2021-09-10 plant biology 10.1101/2021.09.09.459645 medRxiv
Top 0.1%
61.4%
Show abstract

Although growing evidence shows that microRNA (miRNA) regulates plant growth and development, miRNA regulatory networks in plants are not well understood. Current experimental studies cannot characterize miRNA regulatory networks on a large scale. This information gap provides a good opportunity to employ computational methods for global analysis and to generate useful models and hypotheses. To address this opportunity, we collected miRNA-target interactions (MTIs) and used MTIs from Arabidopsis thaliana and Medicago truncatula to predict homologous MTIs in soybeans, resulting in 80,235 soybean MTIs in total. A multi-level iterative bi-clustering method was developed to identify 483 soybean miRNA-target regulatory modules (MTRMs). Furthermore, we collected soybean miRNA expression data and corresponding gene expression data in response to abiotic stresses. By clustering these data, 37 MTRMs related to abiotic stresses were identified including stress-specific MTRMs and shared MTRMs. These MTRMs have gene ontology (GO) enrichment in resistance response, iron transport, positive growth regulation, etc. Our study predicts soybean miRNA-target regulatory modules with high confidence under different stresses, constructs miRNA-GO regulatory networks for MTRMs under different stresses and provides miRNA targeting hypotheses for experimental study. The method can be applied to other biological processes and other plants to elucidate miRNA co-regulation mechanisms.

11
Generate a new crucian carp (Carassius auratus) strain without intermuscular bones by knocking out bmp6

Kuang, Y.; Zheng, X.; Cao, D.; Sun, Z.; Tong, G.; Xu, H.; Yan, T.; Tang, S.; Chen, Z.; Zhang, T.; Zhang, T.; Dong, L.; Yang, X.; Zhou, H.; Guo, W.; Sun, X.

2022-11-28 genetics 10.1101/2022.11.28.518130 medRxiv
Top 0.1%
61.0%
Show abstract

Elimination of intermuscular bones (IMBs) is vital to the aquaculture industry of cyprinids. In our previous study, we characterized bmp6 as essential in the development of IMBs in zebrafish. Knockout of bmp6 results in the absence of IMBs in zebrafish without affecting growth and reproduction. Therefore, we hypothesized that bmp6 could be used to generate new cyprinid strains without IMBs by gene editing. In this study, we established a gene editing strategy for knocking out the two orthologs of bmp6 in diploid crucian carp (Carassius auratus). We obtained an F3 population with both orthologs knocked out, in which no IMBs were detected by bone staining and X-ray, indicating that the new strain without IMBs (named WUCI strain) was successfully generated. Furthermore, we extensively evaluated the performance of the new strain in growth, reproduction, nutrient components, muscle texture and structure, and metabolites in muscle. The results showed that the WUCI strain grew faster than the wild-type crucian carp at 4-month-age. The reproductive performance and flesh quality did not show significant differences between the WUCI strain and wild-type crucian carp. Moreover, the metabolomics analysis suggested that the muscle tissues of the WUCI strain significantly enriched some metabolites belonging to the Thiamine metabolism, Nicotinate and Nicotinamide Metabolism pathway, which plays beneficial effects in anti-aging, anti-oxidant, and anti-radiation damage. In conclusion, we established a strategy to eliminate IMBs in crucian carp and obtain a WUCI strain whose performance was expected compared to the wild-type crucian carp; meanwhile, the WUCI strain enriched some beneficial metabolites to human health in muscle tissue. This study is the first report that a farmed cyprinid strain without IMBs, which could be stably inherited, was obtained worldwide; it provided excellent germplasm for the cyprinids aquaculture industry and a useful molecular tool for eliminating IMBs in other cyprinids.

12
Evolutionary patterns of 64 vertebrate genomes (species) revealed by phylogenomics analysis of protein-coding gene families

Song, J.; Han, X.; Lin, K.

2020-04-01 bioinformatics 10.1101/2020.03.31.017467 medRxiv
Top 0.1%
60.3%
Show abstract

BackgroundRecent studies have demonstrated that phylogenomics is an important basis for answering many fundamental evolutionary questions. With more high-quality whole genome sequences published, more efficient phylogenomics analysis workflows are required urgently. ResultsTo this end and in order to capture putative differences among evolutionary histories of gene families and species, we developed a phylogenomics workflow for gene family classification, gene family tree inference, species tree inference and duplication/loss events dating. Our analysis framework is on the basis of two guiding ideas: 1) gene trees tend to be different from species trees but they influence each other in evolution; 2) different gene families have undergone different evolutionary mechanisms. It has been applied to the genomic data from 64 vertebrates and 5 out-group species. And the results showed high accuracy on species tree inference and few false-positives in duplication events dating. ConclusionsBased on the inferred gene duplication and loss event, only 9[~]16% gene families have duplication retention after a whole genome duplication (WGD) event. A large part of these families have ohnologs from two or three WGDs. Consistent with the previous study results, the gene function of these families are mainly involved in nervous system and signal transduction related biological processes. Specifically, we found that the gene families with ohnologs from the teleost-specific (TS) WGD are enriched in fat metabolism, this result implyng that the retention of such ohnologs might be associated with the environmental status of high concentration of oxygen during that period.

13
ILnc: Prioritizing Long Non-coding RNAs for Pan-cancer Analysis of Immune Cell Infiltration

Li, X.; Yang, C.; Bai, J.; Xie, Y.; Xu, M.; Liu, H.; Shao, T.; Xu, J.; Li, X.

2022-03-13 bioinformatics 10.1101/2022.03.10.483725 medRxiv
Top 0.1%
60.3%
Show abstract

The distribution and extent of immune cell infiltration into solid tumors play pivotal roles in cancer immunology and therapy. Here we introduced an immune long non-coding RNA (lncRNA) signature-based method (ILnc), for estimating the abundance of 14 immune cell types from lncRNA transcriptome data. Performance evaluation through pure immune cell data shows that our lncRNA signature sets can be more accurate than protein-coding gene signatures. We found that lncRNA signatures are significantly enriched to immune functions and pathways, such as immune response and T cell activation. In addition, the expression of these lncRNAs is significantly correlated with expression of marker genes in corresponding immune cells. Application of ILnc in 33 cancer types provides a global view of immune infiltration across cancers and we found that the abundance of most immune cells is significantly associated with patient clinical signatures. Finally, we identified six immune subtypes spanning cancer tissue types which were characterized by differences in immune cell infiltration, homologous recombination deficiency (HRD), expression of immune checkpoint genes, and prognosis. Altogether, these results demonstrate that ILnc is a powerful and exhibits broad utility for cancer researchers to estimate tumor immune infiltration, which will be a valuable tool for precise classification and clinical prediction.

14
Single-cell transcriptomic analysis identifies neocortical developmental differences between human and mouse

Zhou, Z.; Wang, S.; Zhang, D.; Jiang, X.; Li, J.; Gu, Y.; Sun, H.

2020-04-25 bioinformatics 10.1101/2020.04.23.056390 medRxiv
Top 0.1%
60.1%
Show abstract

BackgroundThe specification and differentiation of neocortical projection neurons is a complex process under precise molecular regulation; however, little is known about the similarities and differences in cerebral cortex development between human and mouse at single-cell resolution. ResultsHere, using single-cell RNA-seq (scRNA-seq) data we explore the divergence and conservation of human and mouse cerebral cortex development using 18,446 and 7,610 neocortical cells. Systematic cross-species comparison reveals that the overall transcriptome profile in human cerebral cortex is similar to that in mouse such as cell types and their markers genes. By single-cell trajectories analysis we find human and mouse excitatory neurons have different developmental trajectories of neocortical projection neurons, ligand-receptor interactions and gene expression patterns. Further analysis reveals a refinement of neuron differentiation that occurred in human but not in mouse, suggesting that excitatory neurons in human undergo refined transcriptional states in later development stage. By contrast, for glial cells and inhibitory neurons we detected conserved developmental trajectories in human and mouse. ConclusionsTaken together, our study integrates scRNA-seq data of cerebral cortex development in human and mouse, and uncovers distinct developing models in neocortical projection neurons. The earlier activation of cognition -related genes in human may explain the differences in behavior, learning or memory abilities between the two species.

15
Telomere-to-telomere gap-free and phased genome assembly reveals post-allopolyploidization subgenomic diversification of tobacco centromeres

Chen, W.; Chen, S.; Wang, J.; Meng, D.; Li, J.; Guo, L.

2026-01-22 genomics 10.64898/2026.01.19.700254 medRxiv
Top 0.1%
60.0%
Show abstract

Common tobacco (Nicotiana tabacum) is an economically important crop worldwide whose allotetraploid genome sequence remains incompletely assembled. We report a 4.2Gb telomere-to-telomere gap-free N. tabacum reference genome phased into two subgenomes (subS/subT) , revealing differential evolutionary trajectories between two subgenomes marked by a rapid divergence in paternal T-genome, especially for heterochromatic centromeric regions. To understand the evolutionary dynamics of tobacco centromeres following polyploidization, we characterize the centromere architecture in tetraploid N. tabacum and its diploid progenitors, demonstrating clear subgenomic diversification and repositioning of centromeres during speciation and allopolyploidization. These findings enrich our knowledge on centromere paradox in polyploid genomes.

16
Duck pan-genome reveals two transposon-derived structural variations caused bodyweight enlarging and white plumage phenotype formation during evolution

Wang, K.; Hua, G.; Li, J.; Yang, Y.; Zhang, C.; Yang, L.; Hu, X.; Scheben, A.; Wu, Y.; Gong, P.; Zhang, S.; Fan, Y.; Zeng, T.; Lu, L.; Gong, Y.; Jiang, R.; Sun, G.; Tian, Y.; Kang, X.; Hu, H.; Li, W.

2023-01-30 genetics 10.1101/2023.01.28.526061 medRxiv
Top 0.1%
59.1%
Show abstract

Structural variations (SVs) are a major source of domestication and improvement traits, however SV profiles of duck and their phenotypic impacts largely hidden. We present the first duck pan-genome constructed using five genome assemblies capturing [~]40.98 Mb new sequences. This pan-genome together with high-depth sequencing data ([~]46.5X) identified 101,041 SVs, of which substantial proportions were derived from transposable element (TE) activity. Many TE-derived SVs anchoring in a gene body or regulatory region are linked to ducks domestication and improvement. By combining quantitative genetics with molecular experiments, we dissect how TE-derived SVs change gene expression of IGF2BP1 and generate a novel transcript of MITF, shaping bodyweight and white plumage. In the IGF2BP1 locus, the TE-derived SV explains the largest effect on bodyweight among avian species (27.61% of phenotypic variation). Our findings highlight the importance of using a pan-genome as a reference in genomics studies and explore the roles of TE-derived SVs in trait formation and in livestock breeding.

17
JointMap identified SERPINA3+ chondrocytes as a therapeutic target for osteoarthritis

Wu, H.; Yan, W.; Sun, M.; Fan, Y.; Jia, S.; Wang, J.; Xu, B.; Lu, B.; Du, Y.; Chen, L.; Lin, J.; Cheng, J.; Ao, Y.; Hu, X.

2025-09-15 bioinformatics 10.1101/2025.09.09.675251 medRxiv
Top 0.1%
57.3%
Show abstract

BackgroundOsteoarthritis (OA) is a prevalent degenerative joint disease that affects nearly half of the population aged over 65. Recent single-cell transcriptional studies have identified heterogeneous OA-specific chondrocyte subtypes, yet a unified subtype-classification remains incomplete. MethodsIn this study, we integrated a total of 545,946 cells to develop JointMap, an interspecies atlas. The chondrocyte subtypes were analyzed using machine learning algorithms were validated via human samples, animal models, and in vitro cell culture. We performed KAS-seq and analyzed the spatial characters for JointMap. ResultsBased on clinical diagnosis and cross-omic indexes, we identified a novel OA-elevated chondrocyte subtype, SERPINA3+ chondrocytes. Notably, we uncovered their upregulated chondrogenesis function and intercellular communication patterns as potential therapeutic targets. ConclusionsIn summary, JointMap provides a unified framework for the systematic annotation of chondral tissues across species and helps in the identification of druggable networks for articular diseases.

18
Differentiated genomic footprints and connections inferred from 440 Hmong-Mien genomes suggest their isolation and long-distance migration

He, G.; Chen, J.; Liu, Y.; Hu, R.; Wang, P.; Duan, S.; Tang, R.; Yang, J.; Wang, Z.; Xu, X.; Sun, Y.; Yun, L.; Hu, L.; Yan, J.; Nie, S.; Wei, L.; Liu, C.; Wang, M.

2023-01-17 genetics 10.1101/2023.01.14.523079 medRxiv
Top 0.1%
56.4%
Show abstract

BackgroundThe underrepresentation of Hmong-Mien (HM) people in Asian genomic studies has hindered our comprehensive understanding of population history and human health. South China is an ethnolinguistically diverse region and indigenously settled by ethnolinguistically diverse HM, Austroasiatic (AA), Tai-Kadai (TK), Austronesian (AN), and Sino-Tibetan (ST) people, which is regarded as East Asias initial cradle of biodiversity. However, previous fragmented genetic studies have only presented a fraction of the landscape of genetic diversity in this region, especially the lack of haplotype-based genomic resources. The deep characterization of demographic history and natural-selection-relevant architecture in HM people was necessary. ResultsWe comprehensively reported the population-specific genomic resources and explored the fine-scale genetic structure and adaptative features inferred from the high-density SNP data in 440 individuals from 34 ethnolinguistic populations, including previously unreported She. We identified solid genetic differentiation between inland (Miao/Yao) and coastal (She) southern Chinese HM people, and the latter obtained more gene flow from northern East Asians. Multiple admixture models further confirmed that extensive gene flow from surrounding ST, TK, and AN people entangled in forming the gene pool of coastal southeastern East Asian HM people. Population genetic findings of isolated shared unique ancestral components based on the sharing alleles and haplotypes deconstructed that HM people from Yungui Plateau carried the breadth of genomic diversity and previously unknown genetic features. We identified a direct and recent genetic connection between Chinese and Southeast Asian HM people as they shared the most extended IBD fragments, supporting the long-distance migration hypothesis. Uniparental phylogenetic topology and Network relationship reconstruction found ancient uniparental lineages in southwestern HM people. Finally, the population-specific biological adaptation study identified the shared and differentiated natural-selection signatures among inland and coastal HM people associated with physical features and immune function. The allele frequency spectrum (AFS) of clinical cancer susceptibility alleles and pharmacogenomic genes showed significant differences between HM and northern Chinese people. ConclusionsOur extensive genetic evidence combined with the historic documents supported the view that ancient HM people originated in Yungui regions associated with ancient Three-Miao tribes descended from the ancient Daxi-Qujialing-Shijiahe people. And then, some recently rapidly migrated to Southeast Asia, and some culturally dispersed eastward and mixed respectively with Southeast Asian indigenes, coastal Liangzhu-related ancient populations, and incoming southward Sino-Tibetan people. Generally, complex population migration, admixture, and adaptation history contributed to their specific patterns of non-coding or disease-related genetic variations.

19
Bi-parental graph strategy to represent and analyze hybrid plant genomes

Kong, Q.; Jiang, Y.; Wang, Z.; Wang, Z.; Liu, Y.; Gan, Y.; Liu, H.; Gao, X.; Yang, X.; Song, X.; Liu, H.; Shi, J.

2023-11-29 bioinformatics 10.1101/2023.11.28.568999 medRxiv
Top 0.1%
56.2%
Show abstract

Hybrid plants are universally existed in wild and often exhibit greater performance of complex traits compared with their parents and other selfing plants. This phenomenon, known as heterosis, has been extensively applied in plant breeding for decades. However, the process of decoding hybrid plant genomes has seriously lagged due to the challenges in their genome assembling and the lack of proper methods to further represent and analyze them. Here we report the assembly and analysis of two hybrids: an intraspecific hybrid between two maize inbred lines and an interspecific hybrid between maize and its wild relative teosinte, based on the combination of PacBio High Fidelity (HiFi) sequencing and chromatin conformation capture sequencing data. The haplotypic assemblies are well-phased at chromosomal scale, successfully resolving the complex loci with extensive parental structural variations (SVs). By integrating into a bi-parental genome graph, the haplotypic assemblies can facilitate downstream short-reads based SV calling and allele-specific gene expression analysis, demonstrating outstanding advantages over one single linear genome. Our work provides an entire workflow which hopefully can promote the deciphering of the large numbers of hybrid plant genomes, especially those whose parents are unknown or unavailable and help to understand genome evolution and heterosis.

20
Pan-cancer RNA editing activity reveals complex editing functions and cancer immunotherapy biomarkers

Guo, M.; Xiong, Y.

2026-01-08 bioinformatics 10.64898/2026.01.07.698309 medRxiv
Top 0.1%
55.8%
Show abstract

Pervasive RNA editing in human diversifies the transcriptome and proteome. However, biological functions of most RNA editing sites remain to be uncovered. Here, we developed a computational framework (iPEAPR) to investigate editing functions in various pathways and regulatory elements, by using only summary information of millions of editing sites, in >5,000 cancer samples and in >2,500 normal samples, and large numbers of manually curated functional annotations. We observed heterogeneous editing activities across features. Surprisingly, enhancers were among features showing the highest editing activities. Editing of epithelial-mesenchymal transition was the most significantly associated with patient survival. Moreover, editing associations with cancer stemness, DNA repair deficiency, and tumor immune infiltrations uncovered known and potential regulators of tumor features. We constructed an editing-mediated regulatory network, which revealed new editing modulators, including experimentally validated TNRC6A. Lastly, we demonstrated that EIs can act as biomarkers of both anti-PD1 immunotherapy and other cancer drugs, such as MEK and BRAF inhibitors. Collectively, iPEAPR enabled depicting RNA editing-dependent abnormal functions in cancer and can be applied to more diseases.